Tag
1 article
Z.ai has released GLM-5.3-Flash, a 320B-parameter, 18B-active MoE model with a 1M-token context window and native multimodal capabilities. It features a 3x reduction in attention compute and 4.4x in KV cache usage compared to previous versions.